Papers with visual processing

5 papers
“I’ve Seen Things You People Wouldn’t Believe”: Hallucinating Entities in GuessWhat?! (2021.acl-srw)

Copied to clipboard

Challenge: a problem with natural language generation systems is the generation of tokens that are unrelated to the source input.
Approach: They propose two new models to play the GuessWhat?! referential game . they propose to adapt the best visual processing models available to mitigate this issue .
Outcome: The proposed models generate few hallucinations compared to other models available in the literature.
Picturing Ambiguity: A Visual Twist on the Winograd Schema Challenge (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models have demonstrated remarkable success in tasks like the Winograd Schema Challenge (WSC), showcasing advanced textual common-sense reasoning.
Approach: They propose a framework to isolate models' ability in pronoun disambiguation from other visual processing challenges.
Outcome: The proposed framework isolates the models’ ability in pronoun disambiguation from other visual processing challenges.
GlyphPattern: An Abstract Pattern Recognition for Vision-Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for abstract pattern recognition are easier because they do not involve a natural language description of the pattern.
Approach: They present a dataset that pairs human-written descriptions of visual patterns with three visual presentation styles.
Outcome: The proposed benchmark pairs human-written and human-verified patterns with three visual presentation styles.
Generating Image Descriptions via Sequential Cross-Modal Alignment Guided by Human Gaze (2020.emnlp-main)

Copied to clipboard

Challenge: a long tradition of cognitive studies shows that the interplay between language and vision is complex.
Approach: They propose an approach to image description generation where visual processing is modelled sequentially.
Outcome: The proposed model exploits gaze-driven attention to produce better descriptions . it sheds light on human cognitive processes by comparing different ways of aligning gaze with language production.
LLMs Can Compensate for Deficiencies in Visual Representations (2025.findings-emnlp)

Copied to clipboard

Challenge: a strong language backbone in vision-language models compensates for weak visual features by contextualizing or enriching them.
Approach: They investigate whether strong language backbone compensates for weak visual features . they use CLIP-based vision encoders to perform controlled self-attention ablations .
Outcome: The proposed model compensates for weak visual features by contextualizing or enriching them.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations